Research Synthesis Methods
○ Wiley
Preprints posted in the last 7 days, ranked by how well they match Research Synthesis Methods's content profile, based on 20 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Jaber, A.; Hughes, L.; Cameron, A. C.; Quinn, T. J.
Show abstract
Background: Systematic reviews of clinical prediction models increasingly include studies using artificial intelligence (AI) and machine learning (ML) methods alongside traditional multivariable regression approaches. A previously published Excel tool enabled standardised data extraction using the CHARMS checklist and risk of bias assessment using PROBAST. The recent publication of the PROBAST+AI framework, which distinguishes the assessment of model development quality from the assessment of model evaluation risk of bias and assesses applicability in both parts, necessitates an updated digital instrument applicable across prediction modelling methods. Methods: We updated an open-access Excel tool to incorporate the full PROBAST+AI framework. The updated template incorporates structural separation between assessment of model development quality and model evaluation risk of bias, with applicability assessed in both parts. It also incorporates updated signalling questions, including those addressing methodological issues particularly relevant to AI/ML, and automates the generation of summary tables and graphical displays. Results: The updated tool (CHARMS & PROBAST+AI Template) contains 11 worksheets and supports data extraction and appraisal for up to 30 prediction models. Dedicated, linked worksheets enable separate assessment of model development and model evaluation, with Domain 4 distinguishing among Apparent, Internal, and External evaluation settings. Key updates include dedicated assessments for predictor pre-processing, class imbalance handling and recalibration, data leakage prevention, and replication of the full model development pipeline within resampling procedures. Automated sheets dynamically format tables and summary charts covering PROBAST+AI parts. Conclusions: The CHARMS & PROBAST+AI Excel template provides a standardised, user-friendly, and rigorous digital framework for systematic reviewers appraising traditional statistical and AI-driven clinical prediction models.
Li, S.; Zhang, W.; Xing, X.; Shen, Z.; Wang, Y.; Chen, Z.; Neto, O.; Yu, Y.; Wu, C.; Lin, L.
Show abstract
Background Late-stage cancer incidence is being considered as an earlier endpoint in cancer-screening trials, but its trial-level association with cancer-specific mortality may depend on evidence selection and endpoint harmonization. We evaluated the robustness of this association to source-verified additions. Methods We reconstructed the PubMed corpus underlying a 41-comparison review. Gemini 3.1 Pro Preview was used only to prioritize reports for blinded human reassessment. Reviewers determined eligibility, linked reports from the same trial, harmonized endpoints, and verified comparison-level data. We recalculated unweighted Pearson correlations overall and by cancer type after adding earliest-compatible trial comparisons. Results Among 1209 candidate records, 996 PDFs were assessed. Thirty-three reports absent from the source review were prioritized; 26 were eligible, representing 18 trials, and 8 provided compatible comparisons. Adding these comparisons increased the dataset from 41 to 49 and attenuated the overall correlation from 0.73 (95% confidence interval [CI] = 0.55 to 0.85) to 0.59 (95% CI = 0.37 to 0.75). Updated correlations were 0.49 (95% CI = -0.26 to 0.87) for breast, -0.23 (95% CI = -0.71 to 0.40) for colorectal, and 0.83 (95% CI = 0.54 to 0.95) for lung cancer. One sparse-event comparison influenced the colorectal estimate. Conclusions The overall association was sensitive to evidence composition, and cancer-specific stability varied. Late-stage incidence should be evaluated by cancer type and with prespecified sensitivity analyses for evidence selection and endpoint definitions. Model-assisted prioritization cannot replace human eligibility review, trial reconciliation, and source verification.
Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.
Show abstract
Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.
Sierpe, A.; Yen, R. W.; Milliman, A.; Cady, E.; Ahn, B.; Dade, A. E.; Devito, A. M.; Eckert, B. A.; Gopalan, V. V.; Krasinski, S. C.; MacMartin, M. A.; Musacchio, S. G.; Zhang, J.; Saunders, C. H.
Show abstract
Background Agenda-setting is a fundamental patient-centered communication practice in which a clinician works with a patient to elicit, propose, and organize topics for discussion during a clinical encounter. Various agenda-setting interventions have been developed, including patient-facing tools and clinician training, but their effects have not been systematically evaluated. We aimed to determine the effects of these interventions on encounter, patient, care partner, and clinician outcomes. Methods We searched grey literature and seven databases, including PubMed, from inception through July 2025 for randomized and non-randomized comparative studies of interventions designed to promote or improve clinical visit agenda-setting. Two reviewers independently screened articles and extracted data, with a third reviewer resolving conflicts. We assessed risk of bias using RoB 2 for randomized studies and ROBINS-I for non-randomized studies. We conducted random effects meta-analyses when outcomes were sufficiently comparable, assessed heterogeneity using I2, and rated certainty of evidence using GRADE. Post hoc exploratory subgroup analyses examined study design, adjustment status, and intervention structure. Results Twenty-nine articles describing 22 unique studies met the inclusion criteria, including 13 randomized and nine non-randomized studies. Agenda-setting interventions increased the occurrence of agenda-setting (risk ratio 5.43, 95% confidence interval (CI) 2.06 to 14.28, I2=34.6%) and favored the intervention for concerns addressed when measured as a continuous outcome (standardized mean difference (SMD) 0.37, 95% CI 0.16 to 0.57, I2=65.3%) and overall clinician satisfaction (SMD 0.50, 95% CI 0.23 to 0.78, I2=0.0%). There were no clear differences in the number of concerns raised (mean difference (MD) 0.21, 95% CI -0.19 to 0.61, I2=59.6%), visit duration (MD 0.64 minutes, 95% CI -0.83 to 2.12, I2=51.4%), or overall patient satisfaction (SMD 0.05, 95% CI -0.05 to 0.15, I2=47.0%). Potentially important heterogeneity was present for four of these six outcomes. Post hoc exploratory subgroup analyses did not provide clear evidence that effects varied by study design, adjustment status, or intervention structure. Risk of bias was often high, serious, or critical, and certainty of evidence was low or very low for all pooled outcomes. Conclusions To our knowledge, this is the first comprehensive synthesis of clinical visit agenda-setting interventions. Such interventions may increase the occurrence of agenda-setting and the extent to which patient concerns are addressed without increasing visit length. However, the certainty of evidence was low or very low, and the available evidence does not establish a superior intervention structure.
Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.
Show abstract
We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.
Pinedo-Torres, I.; Taype-Rondan, A.; Vera-Luza, A. A.; Zegarra-Lizana, P. A.; Rojas-Vilca, J. L.; Yovera-Aldana, M.
Show abstract
Objective. To determine the publication rate of abstracts presented at the American Diabetes Association Scientific Sessions and to evaluate the association between statistical significance of study results and subsequent publication. Research Design and Methods. We conducted a retrospective cohort study of abstracts presented at the 2018 American Diabetes Association Scientific Sessions. The primary exposure was study result category (statistically significant vs. non-statistically significant findings), and the primary outcome was publication in an indexed journal within 5 years after conference presentation. Publication status was determined through PubMed/MEDLINE and Scopus searches. Adjusted relative risks (RRs) and 95% CIs were estimated using generalized linear models with Poisson distribution and robust variance. Results. Among 541 included abstracts, 321 (59.3%) were subsequently published in indexed journals. Abstracts reporting statistically significant findings had a higher publication rate than those reporting non-statistically significant findings (61.9% vs. 42.3%; p=0.002). In the adjusted analysis, abstracts with non-statistically significant findings had a lower likelihood of publication compared with those reporting statistically significant findings (adjusted RR 0.71 [95% CI 0.55-0.93]; p=0.013). Conclusions. Approximately four in ten abstracts presented at the ADA Scientific Sessions were not published within 5 years. Abstracts reporting non-statistically significant findings had a lower likelihood of subsequent publication, suggesting persistent publication bias in diabetology research. Future initiatives promoting the interpretation of effect estimates, confidence intervals and clinical relevance, rather than statistical significance alone, may help reduce selective dissemination of evidence
Leonhardt, C.; Birrer, D.; Stauffer, M. F.; Toti, J. M. A.; Gallagher, I. J.; Skipworth, R. J. E.; Laird, B.; Kuemmerli, C.
Show abstract
Importance Non-inferiority trials are becoming increasingly popular in abdominal surgery. The non- inferiority margin is critical in the interpretation and conclusion of these trials. Objective This systematic review aims to assess the methodological and reporting quality of non- inferiority randomized controlled trials in abdominal surgery. Evidence Review Non-inferiority trials were systematically identified by searching Ovid Medline, Embase and the CENTRAL databases from 2006 until December 2025. Randomized controlled trials in adult patients with any type of abdominal surgical intervention in at least one trial arm and a sample size greater than or equal to 100 were eligible for inclusion. The primary outcome was the definition of the non- inferiority margin. Secondary outcomes were the reporting of the non-inferiority margin, the robustness of its estimation, the uncertainty of the point estimate and the adequacy of conclusions. Findings A total of 11 045 trials were identified, of which 101 were eligible, enrolling 44 370 patients. Most trials provided a rationale for the non-inferiority design, while six (5.9%) trials did not. Previous literature was commonly used (n=56; 55.4%), but the non-inferiority margin was most often based on a clinical fixed margin or on historical comparison of the treatment and the active comparator. Based on the margin, investigators tolerated substantially worse outcomes of the treatment compared to the comparator. Conclusions were appropriate based on the confidence interval and the predefined non- inferiority margin in 88 (87.1%) of trials. The clinical judgement of the conclusion was overall adequate. Confidence interval estimations were reported in 16 (15.8%) of trials. Simulation studies were limited by the reporting quality. Conclusions and Relevance Clinical fixed margins are commonly used in abdominal surgery non-inferiority randomized controlled trials, however, substantial shortcomings in reporting limit the interpretability and reproduction of study findings. Based on the findings of this study, guidance on surgical- specific non-inferiority margin definitions is needed.
Alve, S. R.; Rahman, S.; Meem, S. M. A. C.
Show abstract
A dental AI system and a dentist reading the same radiographs form a paired comparison. Published comparative studies often report the two arms separately against a reference standard, leaving the joint pattern of correctness between them unavailable for secondary paired inference. We show what that omission costs. The accuracy difference remains exactly identified; its sampling variance does not, so the report contains the estimate and not its uncertainty. On a study of 282 units, two published accuracies are consistent with 38 distinct joint tables whose confidence intervals differ in width by a factor of 2.5. The consequence is a three-zone decision map rather than a single threshold: differences at or below 1.06 points are non-significant under every compatible table, differences at or above 6.03 points are significant under every compatible table, and in between the published numbers cannot decide. We then show the omission is repairable at negligible cost. One additional integer, the number of units both arms classify correctly, identifies the joint table exactly and restores standard paired inference. For a panel of readers the pairwise dependences must arise from one joint distribution, a constraint that binds once three readers are present; publishing each reader's joint-correct count against a single reference reader cannot widen and may tighten every pairwise bound, and in a 7-arm experiment reduced them by a median of 37% even for pairs excluding that reference. Where the integer was never published we give DentalPair-Cert, an interval with finite-sample coverage uniformly over every admissible within-unit AI-dentist dependence under the independent-sampling-unit model, certified in both the nuisance maximization and the inversion. Across 4,200,000 simulated comparisons an independence analysis falls to 74.5% coverage with 12.2% type-I error; in a purposive sample of 9 recent comparative studies, 1 reported a paired test on discordant units.
Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.
Show abstract
Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.
Dick, M.; Madathil, S.; Patel, A.; Kapoor, H. S.; Sharma, M.; D'Souza, Z.; Hameed, S.; Abu-Samak, M.; Najirad, A.; Dwairi, D.; Radaideh, O.; Nicolau, B.
Show abstract
Objectives: Dentists prescribe approximately one in ten antibiotics worldwide, yet antimicrobial stewardship (AMS) remains underemphasized in dental education. Large language models (LLMs) may support AMS training, but their proficiency and clinical reasoning in this context remain unclear. We evaluated GPT-4o's accuracy and clinical reasoning on dental antibiotic prescribing questions, stratified by question difficulty. Methods: We assembled 125 multiple-choice questions on dental antibiotic prescribing from eight peer-reviewed studies (2017-2023). GPT-4o answered each question and generated a clinical justification. Accuracy was assessed against source-study answer keys and examined across difficulty quartiles. Justifications were evaluated using an adapted 12-axis human-evaluation framework assessing scientific consensus, extent and likelihood of harm, inappropriate and missing content, bias, and both correct and incorrect comprehension, retrieval, and reasoning. Prophylaxis-specific questions were analysed separately. Results: GPT-4o correctly answered 72% of questions. Accuracy remained relatively stable across difficulty quartiles (78%, 78%, 65%, 70%). Experts rated 95.4% of justifications positively across the 12 axes. Comprehension, retrieval, and reasoning each exceeded 96.2% positive ratings. Missing content was the main weakness (7.8%), and 7.1% of justifications showed a moderate-to-severe potential for harm. Performance on prophylaxis-specific questions (98.1%) exceeded non-prophylaxis questions (93.0%). Conclusions: GPT-4o demonstrated moderate-to-high proficiency and clinically defensible reasoning in dental antibiotic prescribing questions. However, residual risks indicate that it is not suitable for unsupervised clinical use but shows potential as a supervised AMS educational tool.
Jafree, D. J.; Sun, M.; Stewart, G. W.; Gishen, F.; Swanton, C.; Motallebzadeh, R.; UCL MB-PhD Outcomes Study Group,
Show abstract
Background: Clinician-scientists translate clinical observation into discovery, trials, and policy, yet this workforce is shrinking across health systems worldwide. Integrated MB-PhD training, pausing medical training to complete a PhD before clinical exposure or specialisation, is one route into this career. We aimed to evaluate the long-term value of MB-PhD training and the barriers to clinical-academic careers these face after graduation. Methods: We evaluated all 131 graduates (29.8% female) who entered the University College London (UCL) MB-PhD programme over a 25-year period (1994-2018). Bibliometric outputs were collated via an inter-linked information system. Concurrently, all 131 graduates were invited to respond to open-ended questions on career benefits and structural barriers; 99 (75.6%) responded, and responses were independently coded into themes, which were then reviewed and confirmed by a Study Group of 107 individuals, including the 91 respondents who agreed to participate further. Results: Graduates produced 5,877 publications (1,141 first-author, 819 corresponding-author), attracting 350,754 citations, with a mean relative citation ratio of 3.30 {+/-} 0.47, approximately three times the field average and sustained across three decades of programme entry. Graduates secured an estimated $157.55 million across 99 grants, released 465 public datasets, and were named investigators on 31 clinical trials across five continents. Among the 99 survey respondents, 49.5% held consultant-grade posts, 72.7% remained research-active, and 25.3% had reached senior academic grade. Open-ended responses were coded into five recurring structural barriers, subsequently confirmed by the Study Group: insufficient protected research time (72.2% of responses), unsupportive training structures and limited career opportunities (36.7%, 24.4% of responses), funding and pay barriers (22.2% of responses), and lack of mentorship or geographical/family constraints (14.4%, 13.3% of responses). Conclusions: Integrated MB-PhD training generates sustained academic productivity and leadership, but structural barriers threaten retention of graduates within clinical-academic careers. Protecting research time, stabilising funding and pay, and reducing geographic instability are needed to retain the clinician-scientists that health systems have already invested in training.
Gabida, M.; Kazonga, E.; Bowa, K.
Show abstract
Abstract Preventable neonatal deaths remain a major public health problem in Zimbabwe, where near-universal antenatal and facility-delivery coverage coexist with a rising neonatal mortality rate. This study evaluated whether institutionalising three core "vital signs" of the community health system (a trained village health worker (VHW) workforce, functional community governance structures, and modified women's and men's participatory learning and action groups) reduces preventable neonatal deaths in Mashonaland West Province. An embedded QUAN (qual) mixed-methods design was used, with a two-arm, parallel-group cluster-randomised controlled trial as the dominant strand. Fifty-two ward-level clusters were randomised 1:1 to the institutionalised community health system package or to standard Ministry of Health and Child Care community services, and 984 pregnant women were enrolled between 1 September 2020 and 31 October 2021, with each mother-infant pair followed to 28 days after delivery, yielding 973 mother-infant pairs for intention-to-treat analysis. The primary outcome was neonatal death within 28 days of life, expressed per 1,000 live births. The primary analysis used a three-level mixed-effects log-binomial regression model with cluster and community-health-worker random intercepts, adjusted for pre-specified covariates. Supervised machine-learning classifiers with leave-one-cluster-out cross-validation, Cox proportional-hazards regression, and multilevel logistic models were fitted as supplementary analyses. An embedded longitudinal process evaluation used key informant interviews and focus group discussions, which were analysed thematically and integrated with the quantitative findings. The neonatal mortality rate was 44.8 per 1,000 live births in the intervention arm versus 110.1 per 1,000 in the control arm. The adjusted risk ratio for neonatal death was 0.43 (95% CI 0.26-0.70; p < 0.001), a 57% relative reduction, with a number needed to treat of 16 mother-infant pairs (95% CI 11-29). Low birthweight (<2,500 g), birth interval under two years, and low community women's literacy were the strongest risk factors, while trained VHWs, functional community governance, early antenatal care, and sustained participatory group attendance were independently protective. The women's and men's groups were protective in a dose-dependent manner, becoming significant at four or more cycles (about 14 meetings) (adjusted odds ratio 0.71; 95% CI 0.60-0.85; p = 0.001). A random forest classifier discriminated against neonatal deaths with a cross-validated area under the curve of 0.842 and a sensitivity of 0.912. Qualitative findings converged with the trial results, identifying male engagement, earlier care-seeking, danger-sign literacy, social-network activation, and community death audits as the behavioural and structural mechanisms of change. Institutionalising the community health system package (trained VHWs, functional governance, early antenatal engagement, and sustained participatory groups) was associated with a substantial reduction in preventable neonatal deaths. The findings suggest that in high-coverage, high-mortality settings, the binding constraint is structural rather than clinical, and that scaling functional community governance and workforce infrastructure in the most disadvantaged communities may accelerate progress toward neonatal survival targets. The principal limitations are a one-year follow-up period, the rarity of neonatal death, and concurrent national programming that only partially reached the control clusters. Trial registration: Pan African Clinical Trials Registry, PACTR202607591142118 (https://pactr.samrc.ac.za/TrialDisplay.aspx?TrialID=PACTR202607591142118); registered retrospectively on 7 July 2026.
Mengi, A.; Bagita-Vangana, M.; Tesine, P.; Laman, M.; Bolnga, J. W.; Ome-Kaius, M.; Kulimbao, J.; Mase, J.; Mal, L. S.; Mnjala, H.; Lee, G.; Cassidy-Seyoum, S. A.; Thriemer, K.; Unger, H. W.
Show abstract
Disseminating study results to participants is an ethical responsibility for researchers but remains uncommon in low- and middle-income countries, and participants preferences for receiving study results are poorly understood. This study examined study result dissemination preferences among pregnant women in a phase III malaria prevention trial in Papua New Guinea (PNG). Participants completed an interviewer-administered questionnaire (survey) assessing their interest in and motivation for receiving trial results and preferences for dissemination methods and content. Associations between participants characteristics and dissemination preferences were explored using multivariable logistic regression analysis. Of 1172 trial participants, 96.0% (1125/1172) completed the survey, and of these 99.6% (1121/1125) wanted to learn about the trial results. The main motivation factors driving participants interest were an acknowledgment of their contribution to research (51.7%; n=579) and a better understanding of the study (45.0%; n=505). Most participants (78.9%; n=884) wanted to learn about the trial findings through written summary and a group meeting with other participants at the nearest clinic (31.1%, n=349). Multivariable regression analysis indicated that participants from rural/peri-urban clinics were more likely to choose non-electronic media dissemination approaches such as a group meeting as compared to urban-dwelling participants. Frequently selected items (>50% of participants) for content included information regarding good results of the study, purpose of the study, medical treatment advances, results specific to me, and how study was conducted. There was heterogenicity in the preference for dissemination content: compared to urban clinics rural clinics are less likely to want to learn about how and why study was conducted and medical and scientific advances. Overall, the majority wanted to learn about trial results, highlighting the importance of integrating dissemination into research activities in PNG. Variation in preferences for mode and content of dissemination between study clinics suggests that dissemination activities could be tailored to local context and preferences.
Lu, Z.; Uddin, S.; Uribe, S.; White, S.; Martins, R. T.; Chau, S.; Mosaddek, A. S. M.; Islam, M. S.; Nahar, N.; Azad, A. K. M.; Hossain, K. M. N.; Choudhury, H. S.; Hasan, K. M. R.; Mosaddek, N.; Rahman, S.; Hossain, M. M.; Sizar, K. M. M. H.; Angione, C.; Lio, P.; Islam, M. T.; Moni, M. A.
Show abstract
Stroke remains a leading cause of mortality and long-term disability worldwide, yet rapid diagnosis is often limited by the shortage of trained radiologists, particularly in resource-constrained settings. Automated analysis of CT imaging offers a potential solution, but existing methods often struggle to achieve clinically generalisable performance while jointly addressing multiple diagnostic tasks. Here we present the Intelligent Integrated Stroke Diagnosis System IISDS, an end-to-end deep learning framework built upon StrokeGNN, a graph-based architecture that integrates 3D contextual feature extraction with U-Net-based 2D lesion segmentation to enable comprehensive stroke analysis from non-contrast CT scans. IISDS performs stroke subtype classification, lesion segmentation and lesion volume estimation within a unified pipeline. To develop and validate the system, we collected and curated BGD-ISD through a collaboration between AI researchers, neurologists, radiologists and clinicians, resulting in a large multi-centre dataset comprising 1,507 CT scans from 597 stroke cases acquired across six hospitals and medical centres in Bangladesh. Across BGD-ISD and multiple publicly available datasets, IISDS achieves state-of-the-art performance on all tasks, improving segmentation accuracy by [≥]0.011 Dice score, reducing lesion volume estimation error by [≥]0.3 average symmetric surface distance (ASSD), and increasing classification performance by [≥]0.018 area under the receiver operating characteristic curve (AUC) compared with existing approaches. These results demonstrate the potential of graph-based deep learning to enable clinically generalisable, automated and scalable stroke diagnosis from CT imaging, supporting rapid clinical decision-making, particularly in healthcare environments with limited access to expert radiological interpretation.
Barzideh, A.; Devasahayam, A. J.; Marzolini, S.; Munce, S.; Sibley, K. M.; Inness, E. L.; Mansfield, A.
Show abstract
Background: Aerobic exercise is recommended during stroke rehabilitation to improve cardiorespiratory fitness and support recovery; however, participation rates remain low. While institutional and system-level barriers have been widely examined, less is known about how individual patient factors influence engagement in aerobic exercise during rehabilitation. Objectives: We aimed to determine whether depressive symptoms, apathy, self-efficacy and outcome expectations for exercise, perceived barriers, or past exercise history were associated with aerobic exercise participation in stroke rehabilitation. Methods: In this prospective cohort sub-study, adults admitted to in- or out-patient stroke rehabilitation at three urban hospitals completed validated questionnaires assessing depressive symptoms, apathy, exercise self-efficacy, outcome expectations for exercise, perceived barriers to being active, and premorbid exercise history. Participants were separated into two groups for analysis: those who completed aerobic exercise during rehabilitation and those who did not. Equivalence testing and between-group comparisons were performed. Results: Sixty-two participants were enrolled; 16 participated in aerobic exercise and 46 did not. Groups were not equivalent on any individual-level factors. Compared to non-participants, those who performed aerobic exercise had significantly higher depressive symptom scores (p=0.0025) and lower self-efficacy for exercise (p=0.0087). Non-participants demonstrated significantly higher apathy (p=0.0007). No significant differences were found for outcome expectations, perceived barriers, or exercise history. Conclusion: Depressive symptoms and lower self-efficacy did not impede aerobic exercise participation during rehabilitation. Increased apathy, however, was associated with non-participation. Findings highlight the need for individually tailored aerobic exercise prescriptions that consider motivational and affective factors to optimize engagement during stroke rehabilitation.
Saba, T. M.; Moudgil-Joshi, J.; Pandit, A. S.; Penn, J.; Mallon, D.; Marcus, H. J.; Grover, P.
Show abstract
Background and Objectives: Recurrence following burr-hole drainage of chronic subdural haematoma (cSDH) occurs in 10-25% of cases, sustained by neovascularisation of the subdural neomembrane supplied by the middle meningeal artery (MMA). MMA embolisation reduces recurrence; whether incidental burr-hole intersection of MMA branches during drainage confers similar benefit is unknown. Methods: We performed a multicentre retrospective cohort study of consecutive adults undergoing burr-hole drainage for cSDH at two UK tertiary neurosurgical centres. Postoperative thin-slice CT was used to classify burr-hole intersection of the underlying MMA groove (no hit, distal-branch hit or main-branch hit) and measure perpendicular burr-hole-to-MMA-groove distance. Co-primary outcomes were radiological recurrence and recurrence requiring intervention. Patient-clustered multivariable logistic regression adjusted for prespecified clinical covariates and treating site. Results: 227 patients (284 operated hemispheres) were included. Radiological recurrence decreased from 34.4% with no branch hit to 22.9% with main-branch intersection, with the gradient confined predominantly to unilateral cSDH. Main-branch intersection was associated with lower adjusted odds of radiological recurrence in unilateral cSDH (adjusted OR 0.30, 95% CI 0.11- 0.81; P = .018), with a similar but non-significant association in the overall cohort (adjusted OR 0.53, 95% CI 0.26-1.07; P = .075). Burr-hole-to-MMA-groove distance demonstrated a more consistent association: in the overall cohort, each 5-mm increase independently increased the odds of radiological recurrence (adjusted OR 1.38, 95% CI 1.04-1.82; P = .025). In unilateral cSDH, each 5-mm increase was independently associated with both radiological recurrence (adjusted OR 1.45, 95% CI 1.03-2.04; P = .034) and recurrence requiring intervention (adjusted OR 1.52, 95% CI 1.05-2.20; P = .027). Conclusion: Main-branch intersection of the middle meningeal artery during routine burr-hole surgery is associated with lower recurrence of unilateral cSDH, while the accompanying burr-hole-to-MMA-groove distance gradient provides biologically plausible support for a dose-response relationship. Together, these findings provide mechanistic rationale for prospective evaluation of intentional neuronavigation-guided MMA targeting (BURR-MMA; NCT07549893).
Sato, J.; Salehjahromi, M.; Zafar, A.; Muneer, A.; Xu, X.; Zhu, E.; Vokes, N. I.; Cascone, T.; Le, X.; Altan, M.; Gardner, E. E.; Sheshadri, A.; Ostrin, E. J.; Salahudeen, A. A.; Li, T.; Merad, M.; Chaudhuri, A. A.; Gerber, D. E.; Kay, F. U.; Godoy, M. C. B.; Carter, B. W.; Shroff, G. S.; Byers, L. A.; Chung, C.; Jaffray, D.; Rice, D.; Liao, Z.; Chang, J. Y.; Vaporciyan, A. A.; Gibbons, D. L.; Wu, C. C.; Heymach, J. V.; Zhang, J.; Wu, J.
Show abstract
Biological aging occurs heterogeneously across individuals and organs. However, current measures of biological age incompletely capture organ-specific differences in health and disease risk. Because chest CT visualizes multiple thoracic organs, it offers an opportunity to quantify structural aging across organ systems. Here, we developed MOSAIC-Age, a framework characterizing eight organ-specific aging clocks on chest CT. The clocks were developed and validated using 9,971 CT scans from CT-RATE and MIDRC, and subsequently locked and applied to two independent prospective cohorts with 35,293 participants from the National Lung Screening Trial and Genetic Epidemiology of COPD study. CT-derived biological age gaps (BAGs) were examined in relation to lifestyle and socioeconomic factors, prevalent comorbidities, incident chronic diseases, and all-cause and cause-specific mortality. Higher BAGs, indicating organs that appeared older on CT than expected for their chronological age, were broadly associated with adverse health characteristics, chronic disease burden, and increased mortality risk. Multiple disease outcomes were associated with aging across several organs, whereas in multivariable analyses including all eight organ-specific BAGs, the remaining associations were more organ specific. A greater number of markedly older-appearing organs and a faster pace of aging were each associated with higher mortality. Together, these findings demonstrate that routine chest CT captures both shared and organ-specific patterns of biological aging and establish CT-derived organ aging as a quantitative imaging biomarker for assessing multi-organ health and long-term disease risk.
Wang, F.; Zhang, Y.-j.; Li, Y.-c.; Li, C.; Yu, H.-F.; Deng, H.-J.; Yu, J.-y.; Xia, H.-m.; Yu, C.; Zhang, Y.; Luo, Z.; Dong, Y.; Pan, X.
Show abstract
BACKGROUND: Cerebral ischemia following subarachnoid hemorrhage (SAH) has traditionally been considered transient because functional alterations of the cerebral microcirculation are thought to be self-limiting. However, we identified a previously unrecognized vasculopathy, perivascular fibrosis of the cerebral microcirculation (PFCM), characterized by excessive type I collagen deposition after SAH. This study investigated the mechanisms underlying PFCM and its subsequent effects on cerebral hemodynamics. METHODS: In vivo SAH was modeled in mice by autologous blood injection, whereas oxygenated hemoglobin (OxyHb) exposure was used to mimic SAH in vitro. Pericyte-deficient mice (Pdgfr{beta}+/-) and pericyte-specific vestigial-like family member 3 (VGLL3) conditional knockout mice (Vgll3{Delta}PC) were generated. Pericyte contractility was measured by nanoindentation and traction force microscopy. Molecular mechanisms were examined using Western blotting, immunofluorescence, CUT&Tag, RNA-seq, transmission electron microscopy, and molecular docking. PFCM, impaired dilation of the cerebral microcirculation, and cerebral autoregulation were assessed by two-photon imaging, transcranial Doppler with continuous blood pressure monitoring, super-resolution ultrasound imaging, and photoacoustic imaging. RESULTS: After SAH, mice developed long-term cerebral autoregulation dysfunction marked by impaired dilation of the cerebral microcirculation, with the abnormality being most evident within the relatively lower blood pressure range. The marked reduction in PFCM in Pdgfr{beta}+/- mice indicated that pericytes were the principal cellular contributors. Mechanistically, OxyHb-induced cytoskeletal remodeling in vitro increased pericyte contractility and promoted nuclear translocation of SAH-upregulated VGLL3. This was followed by increased genomic occupancy, Col1a1 transcriptional activation, and type I collagen deposition. Pericyte-specific VGLL3 knockout abolished PFCM and, consequently, significantly alleviated long-term cerebral autoregulation dysfunction. CONCLUSIONS: Our findings identify PFCM mediated by pericytic VGLL3 as a novel vasculopathy leading to long-term cerebral autoregulation dysfunction after SAH.
Mina, I. K.; Hussain, Y.; Siwy, J.; Catanese, L.; Rupprecht, H.; Beige, J.; Staessen, J. A.; Metzger, J.; Persson, F.; Rossing, P.; Delles, C.; Schanstra, J. P.; Bannaga, A.; Vlahou, A.; Mischak, H.; Arasaradnam, R. P.; Latosinska, A.
Show abstract
Background: Fibrosis, characterised by excessive accumulation of collagen type I (COL1), is a common feature of chronic diseases, including liver diseases (LDs), chronic kidney disease (CKD) and heart failure (HF). COL1 degradation products can be detected in urine by proteomics/ peptidomics analyses and may serve as non-invasive biomarkers of fibrosis. We aimed to identify a common molecular signature of fibrosis across these diseases that may ultimately guide interventions to slow disease progression and prevent organ damage. Methods: Using capillary electrophoresis coupled to mass spectrometry (CE-MS), naturally occurring COL1 degradation products (peptides) in the urine of patients with fibrotic disease, LDs (n=127), CKD (n=263) or HF (n=187), were investigated and compared with the same number of matched controls. Disease-associated COL1 peptides were identified separately for each condition, and peptides showing consistent associations across the three diseases were selected to define a common fibrosis signature. A support vector machine model based on the selected peptides was developed and validated in independent cohorts of patients with LDs (n=110), CKD (n=93), HF (n=32) and controls (n=643). Results: We identified a common fibrotic signature consisting of 50 COL1 degradation products, mainly downregulated in fibrosis. A model based on these peptides achieved a strong performance, with an area under the receiver operating characteristic curve (AUC) of 0.935 (95% confidence interval (CI) 0.917-0.953, p<0.0001) in an external validation cohort comprising pooled disease groups (LDs, CKD, and HF) and controls. Performance was maintained in LDs, CKD and HF, with AUCs of 0.917 (95% CI 0.890-0.944, p<0.0001), 0.951 (95% CI 0.931-0.971, p<0.0001) and 0.950 (95% CI 0.903-0.997, p<0.0001), respectively. The model scores were significantly associated with fibrosis stage in LDs (p=0.0097) and with interstitial fibrosis and tubular atrophy in CKD (p=0.045). Conclusion: A model of urinary COL1 peptides captures a shared collagen degradation signature across organs and diseases, enabling the non-invasive assessment of fibrosis irrespective of its origin. As these peptides exclusively reflect collagen degradation, the findings suggest impaired collagen degradation as a driver in fibrosis. Future clinical studies are warranted to evaluate the utility of this model for early fibrosis detection and earlier implementation of anti-fibrotic interventions.
Masharani, A.; Koreki, A.; Marcelo, M.; Shalfrooshan, K.; Diamos, M.-A.; Santucci, C.; Pillai, K.; Bindman, D.; O'Sullivan, S.; Rugg-Gunn, F.; Sidhu, M.; Yogarajah, M.
Show abstract
Objective: To determine whether paradoxical relief, feeling unusually better after a seizure compared to before it, is more common after functional/dissociative seizures (FDS) than epileptic seizures (ES), quantify its diagnostic accuracy, and explore its relationship with preictal symptoms. Methods: Consecutive patients admitted to a tertiary epilepsy unit for prolonged inpatient EEG monitoring underwent a structured clinical interview on admission, before final multidisciplinary diagnostic classification. Preictal dissociative and autonomic/somatic symptom burden was assessed using items adapted from established questionnaires. Diagnostic classification incorporated clinical history, seizure semiology, video electroencephalography findings, and collateral information. Patients with dual or indeterminate diagnoses were excluded. Associations with paradoxical relief were examined using logistic regression, followed by an exploratory mediation analysis. Results: Of 176 patients assessed, 66 with FDS and 65 with ES were included. Paradoxical relief was reported by 46/66 patients with FDS (69.7%) and 10/65 with ES (15.4%; unadjusted odds ratio [OR] 12.65, 95% confidence interval [CI] 5.57 to 31.09). As a diagnostic signal for FDS, paradoxical relief had 69.7% sensitivity (95% CI 57.1 to 80.4), 84.6% specificity (95% CI 73.5 to 92.4), a positive likelihood ratio of 4.53 (2.51 to 8.19), and a negative likelihood ratio of 0.36 (0.24 to 0.52). FDS diagnosis remained independently associated with paradoxical relief after adjustment (OR 10.59, 95% CI 3.42 to 38.06). In a parallel mediation analysis, dissociative symptom burden showed a significant indirect effect, accounting for 19.5% of the association between diagnostic group and relief, whereas the indirect effect through somatic/autonomic symptom burden was not significant. Significance: Paradoxical relief is substantially more common after FDS than ES and may provide a simple, clinically useful diagnostic signal. Its absence does not exclude FDS, and the finding requires external validation. The association with dissociative symptoms is exploratory and supports prospective investigation of whether relief reflects transient resolution of a disturbed, disembodied preictal state.